Statistical Approach to Modelling of Activity of Phenol’s and its Derivatives against L1210 Leukaemia cells
Sameer Dixit1, Arun K. Sikarwar2
1Department of Chemistry, M. J. P. Govt. Polytechnic College Khandwa, Madhya Paradesh (India)
2Department of Chemistry, Govt. Home Science P. G. College Hoshangabad, Madhya Paradesh (India)
*Corresponding Author E-mail: dixitsameer1@rediffmail.com
ABSTRACT:
A Quantitative Structure-Property relationship (QSPR) model was developed for prediction of Activity of Phenol’s and its congeners against L1210 Leukaemia cells. Murine cell lines such as P388 leukemia, L1210 leukemia, and B16 melanoma, dominated the early years of cancer cell testing both in culture and in mice. In this study we have attempted to develop a multiple linear regression (MLR) model with high accuracy and precision. For this first we prepare several models and then validate them by statistical parameters like Q Factor, PE, PSE, SPRESS etc. and proposed a model which has better prediction power to prediction of Activity against L1210 Leukaemia cells.
KEYWORDS: L1210 Leukaemia cells, QSAR, QSPR, Q Factor, PE, PSE, SPRESS.
INTRODUCTION:
Phenol and its congeners are known to induce caspase-mediated apoptosis activity and cytotoxicity on various cancer cell lines1. Apoptosis, scavenging of radicals, antioxidant, and pro-oxidant characteristics are primarily responsible for the antitumor activities of phenolic compounds. Quantitative structure-activity relationship studies on the cellular apoptosis and cytotoxicity of phenolic compounds have been investigated recently by Selassie and colleagues1 wherein models were developed for various carcinogenic cell lines. Murine cell lines, such as P388 leukemia, L1210 leukemia, and B16 melanoma, dominated the early years of cancer cell testing both in culture and in mice. These quantitative structure-activity relationship models are based on few experimentally obtained physicochemical parameters.
For the justification of models which is the mathematical representation of biological activity and physiochemical properties of phenol derivatives, we use some statistical parameters like Q Factor, PE, PSE, SPRESS etc.
The paper deals with structure-activity relationships of phenols and its derivatives for the development of predictive models from theoretical structural parameters and regression methodology. The quantitative structure-activity relationship studies developed here for the caspase-mediated apoptosis activity and cytotoxicity on murine leukemia cell line (L1210) 2. It is seen that such quantitative structure-activity relationships can provide a better-quality predictive model for the phenolic compounds. The biological activities of phenolic compounds have been calculated based on ridge regression analysis- that clearly gives a better significant correlation compared to the activities predicted.
MATERIAL AND METHODS:
The QSAR equation is linear model which relates variations in biological activity to variations in the values of computed (or measured) properties for a series of molecules3. For the method to work efficiently, the compounds selected to describe the “chemical space” of the experiments (the training set) should be diverse4. A Quantitative Structure/Activity Relationship (QSAR) is the study of the dependence of the chemical structure on an observable experimental property or ‘activity’ over a collection of chemical compounds. Modelling this relationship allows predictions to be made about properties of previously unseen chemical compounds.
We found mostly in many QSAR models single descriptor is not sufficient to express completely of property or activity of given set of compounds. So we use more than one descriptor to achieved goal And this type of analysis known as multiple linear regression analysis ‘MLR’. In order to build linear relationship and test model, the 42 compound data sets was used as training to build model. Finally with the selected eight different descriptors, we will build several linear models using the training data sets and following equations were obtained. Among the generated QSAR models; two models were selected on the basis of various statistical parameters such as squared correlation co-efficient (r2) which is relative measure of quality of fit.
For modeling of activity against L1210 Leukemia cells of phenol derivatives in first model we used eight descriptors Mor29p, Mor20e, Mor04m, Mor23m, FDI, RDF045m, MATS5p, and R3e. There are 42 observations (molecules) are used to built first model for Predicted Activity. By regression Statistics we get correlation coefficient is 0.7395, r2 is 0.5468, Adjusted R Square is 0.4369, and Standard Error is 0.4765 for model-I which described by equation 1.
Predicted Activity = (-2.85923*Mor29p) + (1.586351* Mor20e)+ (0.13532*Mor04m) + (-0.1546 *Mor23m) + (4.718362* FDI)+ (0.080583*RDF045m) + (0.035323*MATS5p) + (0.028945*R3e) -1.99213 ………………….. (1)
For modeling of activity against L1210 Leukemia cells of phenol derivatives in Second model we used eight descriptors Mor04m, Mor23m, FDI, RDF045m, MATS5p, R3e, eHOMO, eLUMO. There are 42 observations (molecules) are used to built second model for Predicted Activity. By regression Statistics we get correlation coefficient is 0.9035, r2 is 0.8164, Adjusted R Square is 0.7719, and Standard Error is 0.3033 for model-II which described by equation 2.
Predicted Activity = (0.037022*Mor04m) + (-1.66835*Mor23m) + (5.73102*FDI) + (0.093915* RDF045m) + (-0.09096* MATS5p) + (0.641258*R3e) + (1.632314*eHOMO) + (-0.43152*eLUMO) -5.16256 ……………….. (2)
Table (i): Observed and Predicted value of Activity against L1210 Leukemia cells using Eq. (2)
|
S. No |
Phenols Derivatives |
Obs. Act. |
Predicted Act. |
Residuals |
Standard Residuals |
|
1 |
4-OCH3 |
4.48 |
4.39535 |
0.08465 |
0.311084 |
|
2 |
4-OC2H5 |
4.64 |
4.75308 |
-0.1131 |
-0.41557 |
|
3 |
4-OC3H7 |
4.85 |
4.72783 |
0.12217 |
0.448976 |
|
4 |
4-OC4H9 |
5.2 |
5.20959 |
-0.0096 |
-0.03523 |
|
5 |
4-OC6H13 |
5.5 |
5.21488 |
0.28512 |
1.047795 |
|
6 |
H |
3.27 |
3.56348 |
-0.2935 |
-1.07849 |
|
7 |
4-NO2 |
3.45 |
3.27682 |
0.17318 |
0.636406 |
|
8 |
4-Cl |
4.29 |
4.17342 |
0.11659 |
0.428438 |
|
9 |
4-I |
3.86 |
3.92119 |
-0.0612 |
-0.22488 |
|
10 |
4-F |
3.83 |
3.6267 |
0.2033 |
0.74712 |
|
11 |
4-NH2 |
5.09 |
5.28721 |
-0.1972 |
-0.72471 |
|
12 |
4-OH |
4.59 |
4.31288 |
0.27712 |
1.018377 |
|
13 |
4-CH3 |
3.85 |
4.01429 |
-0.1643 |
-0.60374 |
|
14 |
4-C2H5 |
3.86 |
4.04927 |
-0.1893 |
-0.69553 |
|
15 |
4-CN |
3.44 |
3.03088 |
0.40912 |
1.503476 |
|
16 |
4-OC6H5 |
4.97 |
4.58831 |
0.38169 |
1.40268 |
|
17 |
4-Br |
4.2 |
3.98502 |
0.21498 |
0.790013 |
|
18 |
4-C (CH3)3 |
4.09 |
3.99527 |
0.09473 |
0.348139 |
|
19 |
3-NO2 |
3.48 |
3.67503 |
-0.195 |
-0.7167 |
|
20 |
3-Cl |
3.87 |
3.61048 |
0.25952 |
0.953727 |
|
21 |
3-C(CH3)3 |
3.88 |
3.72841 |
0.15159 |
0.557074 |
|
22 |
3-CH3 |
3.54 |
3.70712 |
-0.1671 |
-0.61416 |
|
23 |
3-OCH3 |
3.71 |
3.71333 |
-0.0033 |
-0.01222 |
|
24 |
3-C2H5 |
3.71 |
3.78722 |
-0.0772 |
-0.28379 |
|
25 |
3-CN |
3.11 |
2.83795 |
0.27205 |
0.99976 |
|
26 |
3-F |
3.46 |
3.44088 |
0.01912 |
0.07026 |
|
27 |
3-OH |
3.46 |
3.93725 |
-0.4773 |
-1.75385 |
|
28 |
3-NH2 |
4.11 |
4.53405 |
-0.4241 |
-1.55835 |
|
29 |
2-CH3 |
3.52 |
3.61698 |
-0.097 |
-0.35638 |
|
30 |
2-Cl |
3.22 |
3.51751 |
-0.2975 |
-1.09331 |
|
31 |
2-F |
3.2 |
3.54654 |
-0.3465 |
-1.27349 |
|
32 |
2-OCH3 |
3.78 |
4.22852 |
-0.4485 |
-1.64826 |
|
33 |
2-C2H5 |
3.75 |
4.07039 |
-0.3204 |
-1.17741 |
|
34 |
2-OH, 4CH3 |
5.03 |
4.37496 |
0.65504 |
2.407212 |
|
35 |
2-NH2 |
5.16 |
4.79038 |
0.36962 |
1.358309 |
|
36 |
2-CN |
3.3 |
3.00053 |
0.29947 |
1.100522 |
|
37 |
2-NO2 |
3.34 |
3.88898 |
-0.549 |
-2.01744 |
|
38 |
2-Br |
3.44 |
3.56072 |
-0.1207 |
-0.44363 |
|
39 |
2-C (CH3)3 |
4 |
3.84762 |
0.15238 |
0.559981 |
|
40 |
4-C3H7 |
4.04 |
4.19329 |
-0.1533 |
-0.56334 |
|
41 |
4-C4H9 |
4.33 |
4.26609 |
0.06391 |
0.234858 |
|
42 |
4-C5H11 |
4.47 |
4.37032 |
0.09968 |
0.366301 |
RESULT AND DISCUSSION:
We observed that for the models discussed above r value is of order of 0.74 to 0.90. This is better for fitting of values. It is worthy to mention that a model (regression equation) with excellent statistics may not necessary have excellent predictive power. Thus the next step of regression analysis is to examine predictive power of the proposed model this can be easily done by calculating Poglianis quality factor5 Q. This quality factor is defined as the ratio of correlation coefficient (r) to the standard error ‘SE’ (standard deviation ‘sd’). The Q values are high for Biological Activity against L1210 Leukemia cells model-2 (Eq.2) has best predictive powers.
The model-II fulfills the selection criteria’s like correlation coefficient r2 >0.8(0.816381) for in-vivo activity with low standard error of squared correlation coefficient r2_se <0.3(0.0919976) show the relative good fitness of the model and F value too high than tabulated F value show the 99% statistical significance of the regression model. The validation criteria for selection of the model are cross validated squared correlation coefficient q2 >0.8(3.446) for training set. Which show accuracy of the statistical calculation. The cross correlation limit is 0.5 which show inter-pair correlations among the selected descriptors are very low6. This model fulfils all validation criteria with low standard error.
PRESS is a good estimation of the real prediction error of the model, provided that the observations (compounds) were independent. If PRESS is smaller than the sum of squares of the response value (SSY), the model indicates better than chance and can be considered “statistically significant”. The ratio PRESS/SSY can be used also to calculate approximate confidence intervals of prediction of new observations (compounds). To be reasonable QSAR model PRESS/SSY should be smaller than 0.4 and the value of this ratio smaller than 0.1indicates an excellent model. The cross-validated PRESS and SSY as recorded in Table (ii) indicates model-2 (Eq.2) for Biological Activity against L1210 Leukemia cells is a better model compare to model-1(Eq.1) and will give excellent result. If the PRESS value is transformed in a dimensionless term by relatively to the initial sum of squares one obtain Q2 i.e. complement to the fraction of unexplained variance over the total variable. This quantity is also called predictive ability or cross validation correlation coefficient. In such case it is expressed by r2cv or R2cv. It is observed that r2cv <R2. This is found to be the case of Activity model-1 (Eq.1) in present study also.
The magnitude of SPRESS indicates uncertainty in prediction. Like the Standard error of estimation SE the model will be smallest value of SPRESS is considered to have better predictive power. However in our case SPRESS is coincides with SE and is therefore, not useful to explain the predictive power of models. So according to SPRESS model-2 (Eq.2) for Biological Activity against L1210 Leukemia cells is a better model and will give excellent result.
Generally PRESS and Q2 (r2cv or R2cv) have good properties which render than approximate for statistical testing with critical distribution. However, for practical purpose end users the use of square root of PRESS/n seems to be more directly related to the uncertainty of prediction. This ratio is named as Predictive Square Error and symbolized as PSE. This parameter PSE is particularly used when SPRESS coincides with SE. Obviously like SE and SPRESS the lower value of PSE indicates least uncertainty in prediction. The parameter PSE is important because it has the same unity as that of the activity. The PSE values are recorded in Table (ii) indicates that for Biological Activity against L1210 Leukemia cells model-2 (Eq.2) have best predictive powers.
The data presented in Table (ii) indicates that in all the cases of model containing more topological index have R2A is smaller compare to model containing minimum topological index. This means that the added topological descriptor is fair share in given model.
We have also calculated LSE values for proposed models and are recorded in Table (ii) which is in favour of the proposed model. If LSE value is low for property or activity compare to other models than model is better prediction power. i.e. Model has low LSE value has excellent prediction power. The LSE values are low for Biological Activity against L1210 Leukemia cells model-2 (Eq.2) has best predictive powers
It is worthy to mention that a model (regression equation) with excellent statistics may not necessary have excellent predictive power. Thus the next step of regression analysis is to examine predictive power of the proposed model this can be easily done by calculating Poglianis quality factor Q7-10. This quality factor is defined as the ratio of correlation coefficient (r) to the standard error ‘Se’ (standard deviation ‘sd’).
Q= r / sd
Thus the higher the value of r (R) and the lower the value of sd the bigger will be the Q and the better will be the predictive power11-12 of that model.
Table (ii): Cross Validation Parameters for modeling of Activity against L1210 Leukemia cells
|
S. No. |
Statistical Parameters |
Model No.1 |
Model No.2 |
|
1 |
R |
0.739 |
0.904 |
|
2 |
R2 |
0.547 |
0.816 |
|
3 |
SE or Sd |
0.477 |
0.303 |
|
4 |
N |
42 |
42 |
|
5 |
no of Descriptors |
8 |
8 |
|
6 |
PRESS |
7.493 |
3.036 |
|
7 |
SSY |
9.041 |
13.498 |
|
8 |
R2cv |
0.207 |
3.446 |
|
9 |
SPRESS |
0.477 |
0.303 |
|
10 |
PSE |
0.422 |
0.269 |
|
11 |
R2A |
0.437 |
0.772 |
|
12 |
LSE |
7.493 |
3.036 |
|
13 |
PE |
0.610 |
0.583 |
|
14 |
Q=r/sd |
1.552 |
2.979 |
|
15 |
PRESS/SSY |
0.829 |
0.225 |
Figure 1 Correlation of Observed and Predicted value of Activity against L1210 Leukemia cells Using Eq. (2)
CONCLUSION:
By the above discussion and result we conclude that model-2 developed and shown by Eq. (2) is excellent to predict the Biological Activity of phenol derivatives against L1210 Leukemia cells. Statistical approach PRESS and SSY, SPRESS and PSE values support this Model. Higher PE and lower LSE values give it to best predictive power.
Observed value of Biological Activity of phenol derivatives against L1210 Leukemia cells was plotted against and Predicted values Using Eq. (2) shown in Figure 1. The figure clearly indicates there is a significant co-relation between Observed and Predicted values of Biological Activity of phenol derivatives. A maximum molecule shows excellent co-relation for Biological Activity of phenol derivatives against L1210 Leukemia cells.
ACKNOWLEDGEMENT:
Authors are very thankful to Mr. A. P. Sakalle, Principal M. J. P. Govt. Polytechnic college, Khandwa for providing facilities and motivation in the work. The authors would like to thanks Farhan Ahmad Pasha and co-workers for their excellent work on anti-Leukaemia agents which help us to this work.
REFERENCES:
1. Selassie and colleagues, J Med Chem, 48, 7234, (2005)
2. Nandi S. l., Vracko M., Bagchi M. C., Anticancer activity of selected phenolic compounds: QSAR studies using ridge regression and neural networks, Chem Biol Drug Des, 70(5), 424-36, (2007)
3. Haghi, A.K. Methodologies and Applications for Chemoinformatics and Chemical Engineering, IGI Global, 2, (2013)
4. Richon A. B., An Introduction to QSAR Methodology, Network Science Corporation, 5, (2018)
5. Pogliani L., Structure Property Relationships of Amino Acids and Some Dipeptides, Amino Acids, 6,141-153, (1994)
6. Rishikesh V. Antre, Rajesh J. Oswal, et al., QSAR Studies of Substituted Pyrazolone Derivatives as AntiInflammatory Agents, Med chem, 2(6): 126-130 (2012)
7. J. H. van Drie, Curr. Pharm.Des. 2003, 9, 1649.
8. J. H. van Drie, in: Computational Medicinal Chemistry for Drug Discovery, P. Bultinck et al., Eds., Marcel Dekker, 2004.
9. H. Meyer, Arch. Exp. Pathol. Pharmakol. 1999, 42, 109.
10. A. Golbraikh, A. Tropsha, J. Mol.Graph Mod. 2002, 20, 269.
11. A. Tropsha, P. Gramatica, V.J. Gombar, QSAR Comb. Sci., 2003, 22, 69.
12. P. Gramatica, Principles of QSAR models validation: internal and external, QSAR Comb.Sci.2007, 26(5), 694-701.
Received on 27.02.2020 Modified on 24.03.2020
Accepted on 28.04.2020 ©AJRC All right reserved
Asian J. Research Chem. 2020; 13(3):237-240.
DOI: 10.5958/0974-4150.2020.00046.2